evaluation and deployment
Healthcare benchmarks are only as good as their assumptions
In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment. Closing the gap requires making assumptions explicit, testing which assumptions hold, and updating evaluation protocols accordingly. Healthcare LLM benchmarks are one of the main paradigms by which LLMs are evaluated prior to clinical settings. Benchmarks provide a stable goalpost that allow researchers to iterate quickly and measure progress consistently. However, in high-stakes domains like healthcare, that same abstraction becomes a liability.
The Alphabet of Data Science
Artificial Intelligence:: AI is the capability of a machine to imitate intelligent human behavior. BMW, Tesla, Google are using AI for self-driving cars. AI should be used to solve real world tough problems like climate modeling to disease analysis and betterment of humanity. Boosting and Bagging: it is the technique used to generate more accurate models by ensembling multiple models together Crisp-DM: is the cross industry standard process for data mining. It was developed by a consortium of companies like SPSS, Teradata, Daimler and NCR Corporation in 1997 to bring the order in developing analytics models.
Evaluation and Deployment of a People-to-People Recommender in Online Dating
Krzywicki, Alfred (University of New South Wales) | Wobcke, Wayne (University of New South Wales) | Kim, Yang Sok (University of New South Wales) | Cai, Xiongcai (University of New South Wales) | Bain, Michael (University of New South Wales) | Compton, Paul (University of New South Wales) | Mahidadia, Ashesh (University of New South Wales)
This paper reports on the successful deployment of a people-to-people recommender system in a large commercial online dating site. The deployment was the result of thorough evaluation and an online trial of a number of methods, including profile-based, collaborative filtering and hybrid algorithms. Results taken a few months after deployment show that key metrics generally hold their value or show an increase compared to the trial results, and that the recommender system delivered its projected benefits.